Skip to content

feat: account JVM UDF Arrow allocations in Spark task memory - #5027

Open
peterxcli wants to merge 20 commits into
apache:mainfrom
peterxcli:feat/arrow-allocator-as-spark=memory-consumer-for-jvmdispatch
Open

peterxcli wants to merge 20 commits into
apache:mainfrom
peterxcli:feat/arrow-allocator-as-spark=memory-consumer-for-jvmdispatch

Conversation

@peterxcli

@peterxcli peterxcli commented Jul 24, 2026 •

Copy link
Copy Markdown
Member

Which issue does this PR close?

Closes #4174.

Rationale for this change

CometScalaUDFCodegen, currently the only CometUDF implementation, creates output Arrow vectors from the process-wide CometArrowAllocator. These JVM allocations bypass Spark’s task memory accounting.

FFI input vectors only wrap native-owned buffers that are already accounted by Comet, so they must not be charged again. Output buffers, however, are created on the JVM and may remain alive through the FFI release callback after the Java vector closes.

What changes are included in this PR?

  • Add a per-task Arrow child allocator backed by a non-spillable Spark MemoryConsumer.
  • Pass the allocator through the generic CometUDF interface and codegen output path.
  • Keep imported native buffers on the root allocator to avoid double-accounting.
  • Charge output buffers to the Spark task while the UDF holds them; at export, buffer accounting moves to the root allocator along with native ownership, so native operators that retain the buffers (hash join builds, shuffle) take the only Spark reservation instead of charging the same physical buffers twice.
  • Accounting requires off-heap Tungsten memory (spark.memory.offHeap.enabled); with on-heap Tungsten memory the off-heap Arrow buffers have no matching Spark pool and are tracked by the allocator but not charged.
  • Interface note: CometUDF.evaluate gains an allocator parameter, so existing implementations need updating. The doc also now states that evaluate may be called concurrently from multiple Tokio workers within one task — this was already true before this PR (CometScalaUDFCodegen has synchronized its body since feat(experimental): ScalaUDF and Java UDF support via Janino codegen #4267 for exactly this reason); the previous "at most one thread" wording was incorrect.

How are these changes tested?

Added CometUdfBridgeTest covering allocation accounting, export ownership transfer, task completion, straggler evaluations racing teardown, and final cleanup. The focused test passes with Spark 4.1 and Spark 3.5/Scala 2.12, and CometCodegenSuite passes end-to-end.

Also ran few scala script with spark shell to verify:

Check main branch
Small output task peak 0 262,144 bytes
32 MiB output task peak 0 50,397,184 bytes
Peak increase 0 50,135,040 bytes
256 MiB batch under 256 MiB pool Succeeds, bypassing Spark Rejected by TaskMemoryManager
Same workload with batch size 32 — Succeeds, 268,435,456 bytes processed

@peterxcli

Copy link
Copy Markdown
Member Author

@sunchao would you like to take a look at this? TIA!

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two P2 findings in task-allocator concurrency, detailed inline. Validated with targeted JVM reproductions against this exact head and Arrow 18.3.0. The native producer/cleanup paths were source-traced, not run end-to-end.

Comment thread spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java Outdated
Comment thread spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java Outdated
@peterxcli
peterxcli requested a review from sunchao August 23, 2026 12:53
@sunchao

sunchao commented Aug 23, 2026

Copy link
Copy Markdown
Member

@peterxcli can you check the CI failures?

@peterxcli

Copy link
Copy Markdown
Member Author

@peterxcli can you check the CI failures?

@sunchao I opened a fix for it

#5439

@sunchao

sunchao commented Aug 23, 2026

Copy link
Copy Markdown
Member

Ah thanks! Approved the PR and will merge soon

Comment thread spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java
peterxcli and others added 2 commits August 27, 2026 23:07
Native operators that retain an exported UDF result (hash join builds,
shuffle buffers) take their own Spark reservations for those buffers, so
keeping the JVM task charge until the FFI release charged the same
physical buffers twice and could fail a native reservation that fit the
pool.

At export, transfer the result's buffer accounting from the task
allocator to the root allocator (zero-copy; FFI addresses unchanged) and
release the corresponding Spark charge. The UDF is still charged while
it holds the memory; once native execution owns the buffers, the only
Spark charge is whatever a retaining native operator reserves. Neither
an ownership transfer nor the eventual FFI release fires the task
allocator's AllocationListener, so the explicit release cannot
double-free, and getAccountedSize() being non-zero only on the owning
ledger keeps pass-through inputs and re-exported chunks at zero.

The C schema is exported from the result's own Field: a transferred
complex vector rebuilds child names from runtime data vectors, which
Arrow hardcodes to "$data$". The C array carries no names, so it comes
from the transferred vector.

Exported buffers no longer pin TaskState past task completion, so the
deferred-release retention (and the task-attempt-ID reuse concern it
covered) is gone.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@andygrove

Copy link
Copy Markdown
Member

Note on this review: this was generated by an LLM (Claude Code) at my request while I worked through a review backlog. I have not verified the individual findings myself. Please treat everything below as suggestions to evaluate rather than as authoritative review feedback, and push back on anything that is wrong or already handled.

Getting JVM UDF output vectors onto Spark's task accounting is worth doing, and the distinction between imported native buffers (already accounted, keep on the root allocator) and JVM-created output buffers (charge to the task) is the right one. The retention logic that keeps the allocator alive until the last FFI release, while dropping Spark accounting at task completion, is carefully thought through.

Four things.

The threading contract for CometUDF changed

The class doc goes from:

At any instant at most one thread is inside evaluate() for a given taskAttemptId.

to:

Calls for a task may arrive concurrently from different Tokio workers. Implementations with mutable state are responsible for synchronizing evaluate().

That is a breaking change to the contract that user-written CometUDF implementations were told they could rely on. Anyone who wrote a UDF holding mutable per-task state, which the old wording explicitly invited, now has a data race.

Was the old guarantee wrong all along, or does this PR change when concurrent calls can happen? Either way this deserves its own line in the description and probably a migration-guide note, because it is a much bigger deal for users than the memory accounting is. Right now it reads like an incidental doc edit.

Lock ordering deserves to be written down

onPreAllocation takes the TaskMemoryManager monitor, then the TaskState monitor, then calls back into TaskMemoryManager.acquireExecutionMemory. taskCompleted and onRelease take the TaskState monitor and then call into MemoryConsumer.freeMemory, which reaches the MemoryManager monitor. Spark's own acquireExecutionMemory can call spill() on other consumers while holding the TaskMemoryManager monitor, and an Arrow buffer release from such a path would re-enter onRelease.

I worked through it and did not find an actual inversion, but it took a while and I am not certain. Could you add a short comment naming the three locks and the order they must always be taken in? Anything that acquires a Spark memory lock from inside an Arrow allocation listener callback deserves that much.

On-heap Tungsten mode silently does nothing

this.consumer = taskMemoryManager.getTungstenMemoryMode() == MemoryMode.OFF_HEAP
    ? new TaskMemoryConsumer(taskMemoryManager) : null;

With on-heap Tungsten memory the whole accounting path is disabled. That is probably correct, since Arrow's buffers are off-heap and charging them to Spark's on-heap pool would be wrong. But it is not stated anywhere, and a user running on-heap who reads the release notes will believe their UDF memory is accounted. Could the class doc say so explicitly, and ideally log once at debug level when accounting is skipped?

The registration-order dependency

registerTask has to be called before CometExecIterator registers its own completion listener, because Spark runs listeners in reverse order. That is documented in the javadoc, which is good, but it is enforced by nothing. If someone reorders those two calls the failure is a use-after-free of an allocator with buffers still exported, which will not reproduce reliably.

Is there a way to make that structural rather than conventional, for example having CometExecIterator obtain the TaskState and register both listeners itself? If not, is one of the tests in CometUdfBridgeTest specifically asserting the ordering?

- Document the lock order (TaskMemoryManager monitor -> TaskState
  monitor -> MemoryManager monitor) on TaskState, verified against
  Spark 3.5.8 and 4.1.3: acquireExecutionMemory holds the TMM monitor
  across spills, releaseExecutionMemory never takes it, so every
  release-side path (TaskState -> MemoryManager) is a suffix of the
  acquire-side order and no inversion exists.

- State explicitly that on-heap Tungsten memory disables Spark
  accounting for UDF Arrow buffers (they are off-heap, so there is no
  matching pool), and log that at debug level once per task.

- Ground the CometUDF concurrency contract in its actual causes:
  DataFusion pipelining via spawned Tokio tasks (HashJoinExec
  OnceAsync) and multiple prefetching native plans per Spark task.
  CometScalaUDFCodegen has synchronized evaluate for this reason since
  its introduction; the old "at most one thread" trait doc was already
  contradicted on main.

- Make completion-listener registration structural rather than
  order-dependent: register outside computeIfAbsent so a straggler
  evaluation racing task teardown recreates state safely instead of
  re-entering the map from its own mapping function, and clarify that
  listener ordering is for orderly shutdown, not memory safety. New
  test covers stragglers both during and after teardown.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@peterxcli

Copy link
Copy Markdown
Member Author

@andygrove Thanks — evaluated all four; addressed in 1dc8fc7.

(1) Threading contract: this is a doc correction, not a behavior change. The "at most one thread" wording was already contradicted on main: CometScalaUDFCodegen — the only CometUDF implementation — has run evaluate under this.synchronized since its introduction in #4267, with a comment explaining that HashJoinExec pipelines build/probe via OnceAsync (tokio::spawn), so multiple Tokio workers can call one task's dispatcher. Two more sources in jni_api.rs: the scan-free path drives each plan from a spawned prefetching Tokio task, and one Spark task can hold several such plans concurrently. Since this PR also changes the evaluate signature (allocator parameter), any out-of-tree UDF must be revisited regardless. I've grounded the trait doc in these concrete causes and will call it out in the PR description.

(2) Lock ordering: added a javadoc on TaskState naming the order — TaskMemoryManager monitor → TaskState monitor → MemoryManager monitor — verified against Spark 3.5.8 and 4.1.3 sources: acquireExecutionMemory holds the TMM monitor across spills (so a spill releasing Arrow buffers re-enters onRelease in the same TMM→TaskState order), and releaseExecutionMemory never takes the TMM monitor, so every release-side path is a suffix of the acquire-side order.

(3) On-heap mode: agreed — the behavior is correct (Arrow buffers are off-heap; there is no matching Spark pool), but it is now stated in the TaskState javadoc and the CometUDF contract, with a per-task debug log when accounting is skipped.

(4) Registration order: the failure mode is not use-after-free — the allocator only closes once in-flight evaluations finish and it holds no memory, and post-completion evaluation is rejected. But investigating this surfaced a real edge: a straggler evaluate() after state removal re-created state via computeIfAbsent, whose mapping function registered a completion listener Spark can invoke immediately — re-entering TASKS.remove from inside computeIfAbsent. Fixed by registering the listener outside computeIfAbsent, which makes the ordering structural (a graceful-shutdown nicety rather than a safety requirement), and added stragglerEvaluationsAroundTaskCompletionAreSafe covering evaluations racing and following task completion (Spark's deferred-listener-drain semantics verified identical in 3.5.8 and 4.1.3).

@peterxcli
peterxcli requested a review from sunchao August 27, 2026 18:10

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two P2 findings in output-buffer accounting, detailed inline. The empty-output issue was reproduced through a native SQL query; the shared-scratch issue was reproduced through the public bridge with a custom CometUDF.

Comment thread spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java Outdated
Comment thread spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java
…F export charge

chargedOutputSize computed the Spark charge to release at export from
getBuffers(false), which omits the allocated buffers of zero-length
children (an empty list's data vector keeps its allocated capacity)
even though TransferPair.transfer() moves those chunks to the root
allocator; the uncounted bytes stayed reserved against the task until
completion. It now enumerates each vector's physical field buffers
recursively, preserving per-ledger deduplication.

It also released the full chunk charge for buffers shared with a
retained scratch vector (an aligned splitAndTransfer slice shares the
scratch ledger). When native execution released the FFI result, Arrow
handed chunk ownership back to the surviving scratch ledger without a
listener callback, and the eventual scratch close fired onRelease and
freed the same charge a second time. A chunk is now counted only when
every live reference to its ledger comes from the result tree; shared
chunks keep their Spark charge and release it exactly once through
onRelease, or wholesale at task completion if scratch closes while
native still holds the buffers.

Both regression tests fail against the previous logic: the empty-child
test strands 32768 bytes of task charge at export, and the shared
scratch test observes the charge dropped while a task ledger still
references the chunk.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@peterxcli
peterxcli requested a review from sunchao August 28, 2026 06:48
@andygrove andygrove added enhancement New feature or request area:udf labels Sep 6, 2026
@andygrove andygrove added the area:memory Memory pools, reservations, OOM handling label Sep 6, 2026
@github-actions github-actions Bot added the area:expressions Expression evaluation label Sep 9, 2026
@andygrove

Copy link
Copy Markdown
Member

Triage note: #5998 overlaps with this. It attaches an AllocationListener to the process-wide CometArrowAllocator root so that every JVM Arrow allocation — FFI export, broadcast coalescing, CometSparkToColumnarExec, the cached batch serializer, codegen output and the Python runner — is reserved from Spark's TaskMemoryManager in fixed blocks. That one is mine, and it is the first item of #5997.

They are not the same change. This PR cuts a per-task child allocator for the UDF output path specifically and charges it to a non-spillable MemoryConsumer, which is more precise on that path than a root-level listener, and the CometUDF interface change has no counterpart in #5998. But both charge the same output buffers and both touch CometUdfBridge.java, so landing them unchanged would double-count. Could we work out which layer owns the UDF output buffers before either goes further?

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: JVM UDF output allocations bypassed Spark task memory accounting.
  • Design approach: Charge a task-scoped Arrow allocator through a non-spillable Spark consumer, then transfer output ownership at export. The transfer preserves shared buffers without copying their contents.
  • Correctness / compatibility analysis: Found one introduced P1 compilation failure detailed below. Reviewed allocation, release and completion behavior against Spark 3.4.3, 3.5.9, 4.0.4, 4.1.3 and 4.2.0 sources and Arrow 18.3.0. Earlier review fixes are present, and no remaining existing P1/P2 concern was substantiated.
  • Key design decisions: Task-context identity separates task lifetimes. In-flight evaluation guards protect teardown. Recursive buffer accounting handles empty children and shared ledgers, keeping ownership logic within the bridge.
  • Implementation sketch: CometExecIterator registers task state, codegen receives the allocator through CometUDF.evaluate, and the bridge transfers exported results to the root allocator.
  • Behavioral changes worth calling out: Off-heap UDF allocations now face Spark admission control. On-heap mode skips Spark charging. Custom implementations must adopt the new allocator parameter. Runtime performance was not measured.
  • Suggested improvements: Update both benchmark callers for the new signature and rerun the affected build and tests.

Reviewed the entire 10-file diff from 36146a87bf9ca9ca9e211b4372628ed2f9d8c8c6 to dc117b0135659137591e26e84edd87ee1a387e03. Confirmed the PR is non-draft. Read existing reviews, issue comments, inline comments and threads. Routed skill: review-comet-pr. Checked sibling skill scopes; none applies to this allocator and lifecycle change.

Exact-head CI: All four Linux Spark 4.1 test groups and both TPC verification jobs fail during test compilation, making Required Checks fail. JVM/native builds, Rust tests and lint checks pass. The logs identify the merge of the requested head and base.

Validation limits: The local focused test attempt stopped before compilation because Maven Central access was unavailable. No local runtime tests or end-to-end queries ran. The finding is independently confirmed by exact-head CI logs and the unchanged benchmark callers. Project code remains unchanged.


override def evaluate(inputs: Array[ValueVector], numRows: Int): ValueVector = {
override def evaluate(
allocator: BufferAllocator,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Update both benchmark callers when adding the allocator parameter. CometTimeExtractBenchmark.scala still calls dispatcher.evaluate(inputs, size) at lines 122 and 149. With the default Spark 4.1 profile, test compilation now fails instead of compiling those calls, preventing every Linux JVM test group and both TPC verification jobs from running. Please pass CometArrowAllocator at both sites, as the updated direct call in CometCodegenSuite does, then rerun test compilation and the focused suites.

Evidence: Exact-head CI run 35170049560 checks out the merge of dc117b0 into 36146a8. The expressions job reports not enough arguments for method evaluate and Unspecified value parameter numRows at CometTimeExtractBenchmark.scala:122 and :149, then fails scala-maven-plugin:4.9.6:testCompile. The same errors occur in the exec, scans, shuffle, TPC-H and TPC-DS jobs. Source comparison confirms the callers are unchanged while the two-argument method was replaced by this three-argument signature. https://github.com/apache/datafusion-comet/actions/runs/35170049560/job/105040966472

# Conflicts:
#	spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java
@andygrove

Copy link
Copy Markdown
Member

I closed #5998 after finding that charging JVM Arrow memory to Spark, and refusing on a short grant, turns native spills into task failures. This PR refuses the same way: onPreAllocation throws OutOfMemoryException when acquireMemory comes back short.

NativeMemoryConsumer.spill returns 0, and native operators keep reserving until try_grow fails. So by the time a task is under pressure, native has already filled its share. The JVM allocation for the next batch asks just before native does, and that allocation is the one refused. In #5998's repro, at local[4] with 128m off-heap, a sort and a native shuffle over a Comet-cached table failed with Unable to reserve ... got 0 in all four tasks. Both spilled normally with the refusal switched off. CI can't see this, because CometTestBase runs a 2 GiB pool.

UDF output goes straight into native operators that could have spilled instead, so I'd expect the same failure here. Could the consumer record the allocation without refusing it? Or refuse only when nothing native in the task can spill? For the rest of the JVM Arrow memory, I ended up with #6250: it reports the figures in the memory usage log and doesn't gate on them.

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: JVM UDF output allocations bypassed Spark task memory accounting.
  • Design approach: Introduce a task-scoped Arrow allocator backed by a non-spillable Spark consumer, then transfer output ownership at export.
  • Correctness / compatibility analysis: No additional introduced P1/P2 issues found within this review. The earlier benchmark compilation blocker is fixed. The existing memory-pressure concern remains substantiated and unresolved, as detailed below.
  • Key design decisions: Task-context identity separates task lifetimes. In-flight evaluation guards protect teardown. Recursive ledger accounting handles empty children and shared buffers. Ownership logic stays localized in the bridge.
  • Implementation sketch: CometExecIterator registers task state, CometUDF.evaluate receives the allocator, and codegen allocates outputs through it. Export transfers buffers to the root allocator and releases exclusive output charges.
  • Behavioral changes worth calling out: Off-heap UDF allocations now face Spark admission control. On-heap mode skips Spark charging. Custom UDF implementations must adopt the allocator parameter. Fresh codegen outputs retain zero-copy export. Runtime overhead was not benchmarked.
  • Suggested improvements: Resolve the existing interaction between UDF admission control and native spilling before merge, with a small-pool UDF-to-sort regression test.

Reviewed the entire 12-file diff from 81f2574d5e92028d6a0c8aebd39e6e0b5b4c2fd2 to dd24a38d428fd0cddc1a0cf1ce1601a4be91a3a3. Confirmed the PR is non-draft. Read AGENTS.md, existing reviews, issue comments, inline comments and review threads. Routed skills: review-comet-pr, review-comet-memory-pr, and review-comet-ffi-pr. Compared relevant behavior against Spark 3.4.3, 3.5.9, 4.0.4, 4.1.3 and 4.2.0 sources, plus Arrow 18.3.0.

The existing memory-pressure concern remains a P2 blocker. An exact-source component probe reserved 67,104,768 bytes through CometTaskMemoryManager in a 64 MiB pool. A 1,024-row UDF requiring 8,192 Arrow bytes succeeded with the root-allocator control but failed with the task allocator. Partial grants were correctly returned, and the same allocation succeeded after releasing native reservations. Source tracing confirms NativeMemoryConsumer.spill() returns zero, while DataFusion’s sorter propagates an input error before reaching insert_batch and its spill-and-retry path. This can turn spillable workloads into task failures. It is already discussed, so no duplicate finding is added.

Exact-head CI: Run 36243579843 is green, including Required Checks, all four Linux Spark 4.1 test groups, native/Rust checks and TPC-H/TPC-DS verification. Logs confirm the merge includes the requested head. Spark SQL, macOS and other Spark runtime profiles were skipped.

Validation limits: All six CometUdfBridgeTest tests passed locally in an isolated build of the exact bridge sources with the original allocator declarations, Spark 4.1.3, Arrow 18.3.0 and JDK 21. The pressure reproduction was a JVM component probe, not an end-to-end native sort query. No full local project build or Spark SQL suite ran. Project files remain unchanged.

… grants

The per-task UDF allocator's listener threw OutOfMemoryException from
onPreAllocation whenever Spark's acquireMemory granted less than requested.
Native operators reserve through a consumer whose spill returns 0 and spill
only when their own try_grow fails, so under pressure native execution has
already filled the pool and the next UDF output allocation is the one
refused, failing the task where a downstream native sort or shuffle could
have spilled.

The listener now charges whatever Spark grants and carries the rest as a
shortfall, like the native pools' grow/overcommit. Releases (buffer close,
failed allocation, export to native) repay the shortfall before returning
anything to Spark, so Spark stays charged for min(its grant, outstanding
bytes) and is never handed back more than it granted. The consumer also
clamps every release to what it still holds, which keeps Spark's books
right when Arrow moves chunks between allocators without a listener
callback. Task completion returns exactly the outstanding grant. The only
remaining refusal is an allocation after task completion.

Tests: the blocking-release test now requires the pending allocation to
succeed; two JUnit tests reproduce the small-pool probe (native consumer
holding all but 4 KiB of a 64 MiB pool) through the public bridge and on
the release and completion paths; CometCodegenMemoryPressureSuite runs
regexp_replace (codegen dispatch) into a native sort in a 16 MiB pool and
asserts the sort spills. All fail on the previous head with "Failed to
acquire ... JVM UDF bytes from Spark TaskMemoryManager".
@andygrove

Copy link
Copy Markdown
Member

This is a light fully automated review since there are so many PRs open.

  1. chargedOutputSize (spark/src/main/java/org/apache/comet/udf/CometUdfBridge.java:395) assumes every byte a task-allocator ledger owns was charged in onPreAllocation. But Arrow's transferOwnership moves ownership through forceAllocate, which never calls the listener, and imported inputs are owned by a ledger on CometArrowImportAllocator. So if a UDF returns an input moved into allocator with getTransferPair(allocator) and transfer(), or an aligned splitAndTransfer slice of one (the comment at line 284 already expects input slices), the ref-count check still passes and releaseExportedCharge frees bytes this consumer never acquired. Spark releases from the task-wide total, so that comes out of the task's native reservations, and accounted goes negative. Closing a transfer()-ed input inside the UDF does the same through onRelease. Could releases be capped at what this consumer holds, with a test where another consumer holds memory while the UDF returns a transferred input?

  2. taskCompleted (CometUdfBridge.java:614) drops UDF instances without closing anything, and closeIfIdle waits for the allocator to be empty. CometUDF.scala:52 still invites scratch buffers in fields, and line 38 now asks for temporary buffers to come from allocator. A UDF holding one leaves the allocator non-empty for good. Arrow's root keeps every unclosed child in childAllocators, so each such task pins its TaskState, TaskContext and TaskMemoryManager for the executor's lifetime, where before only the off-heap bytes leaked. Would a close() hook on CometUDF, called once the task is complete and nothing is in flight, make sense?

  3. docs/source/contributor-guide/memory_management.md:133 still routes CometBatchKernelCodegenOutput to CometArrowAllocator and "accounted by nobody", the child allocator list at line 112 has no per-task UDF allocator, and line 167 calls CometTaskMemoryManager the one place Comet acts as a Spark MemoryConsumer. Could this PR update those?

@sunchao sunchao left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary

  • Prior state and problem: JVM UDF Arrow allocations bypassed Spark task memory accounting.
  • Design approach: Introduce a task-scoped allocator backed by a non-spillable Spark consumer. Record short grants without rejecting allocations, allowing downstream native operators to spill.
  • Correctness / compatibility analysis: No additional introduced P1/P2 issues found within this review. Earlier compilation and memory-pressure blockers are fixed. The existing scratch-buffer retention concern remains a reproducible P2 blocker.
  • Key design decisions: Task-context identity separates task lifetimes, in-flight guards protect teardown, and recursive ledger inspection handles empty children and shared buffers. Ownership logic remains centralized in the bridge.
  • Implementation sketch: CometExecIterator registers task state, CometUDF.evaluate receives the allocator, and export transfers output ownership to the root allocator without copying buffer contents.
  • Behavioral changes worth calling out: Custom UDF implementations must adopt the allocator parameter. On-heap mode skips Spark charging. Short grants preserve execution but do not bound physical memory. Runtime accounting overhead was not benchmarked.
  • Suggested improvements: Resolve the existing scratch-lifecycle concern by providing cleanup after in-flight evaluations finish, with a repeated-task regression test.

The existing scratch-buffer concern remains unresolved. A custom UDF retaining an 8,192-byte scratch buffer from the supplied allocator left one additional task state after each completed task. Three sequential tasks retained three child allocators and three completed task contexts. taskCompleted() clears UDF instances at CometUdfBridge.java:652, while closeIfIdle() requires an empty allocator at line 677. Consequently, the pre-existing scratch-buffer leak now also retains task state and its memory manager indefinitely. This is already reported, so no duplicate finding is added.

Reviewed the entire 16-file diff from 81f2574d5e92028d6a0c8aebd39e6e0b5b4c2fd2 to e8ed305d28643fe0fd13071d57e9ec4bf24351a5. Confirmed the PR remains non-draft. Read AGENTS.md, existing reviews, issue comments, inline comments and threads. Routed skills: review-comet-pr, review-comet-ffi-pr, and review-comet-memory-pr. Compared relevant lifecycle and accounting semantics against Spark 3.4.3, 3.5.9, 4.0.4, 4.1.3 and 4.2.0 sources, plus Arrow 18.3.0.

Exact-head CI: Run 36468098758 passed, including all four Linux Spark 4.1 test groups, native/Rust checks, TPC verification and the new memory-pressure regression. Logs confirm the tested merge includes the requested head. A duplicate cancelled run produced a failed Required Checks result. Spark SQL, macOS and other Spark runtime profiles were skipped.

Validation limits: All eight CometUdfBridgeTest tests passed in a disposable component build using exact checkout sources and the original allocator declarations, with Spark 4.1.3, Arrow 18.3.0 and JDK 21. The scratch reproduction exercised the public JVM bridge and Arrow FFI, not a native query. No full local project build or Spark SQL suite ran. Project code and GitHub state remain unchanged.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:expressions Expression evaluation area:memory Memory pools, reservations, OOM handling area:udf enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Register CometArrowAllocator as a Spark MemoryConsumer for JVM-UDF dispatch

3 participants